Learning without labels
Supervised learning, which you covered in the last two lessons, requires labelled examples. Someone has to go through the data and write down the correct answer for each row. That is expensive, slow, and sometimes impossible. You cannot label every customer, every web page, every image on the internet.
Has labels. The model learns from input-output pairs. You know the correct answer for every training example.
image → cat or dog
house features → price
No labels. The model finds patterns in raw data without being told what to look for. The structure emerges from the data itself.
articles → topics
genes → expression patterns
Unsupervised learning is what you turn to when you have data but no answers. When a streaming service wants to understand its user base without manually categorising millions of accounts. When a biologist wants to see which genes behave similarly without having any prior theory about the groupings. When a retailer wants to understand natural buying patterns before designing targeted promotions.
Clustering is the most widely used unsupervised technique. The goal is simple: group data points so that points in the same cluster are similar to each other and different from points in other clusters.
You have just inherited a huge box of old photographs from a relative. No names, no dates, no labels. But as you spread them across a table, you start to notice things. These 30 photos all seem to feature the same family at a beach house. These 20 are all winter holidays. These 15 all look like work events. You are clustering: grouping similar things together without anyone telling you how many groups to make or what to call them. K-Means does the same thing with numbers.
K-Means: the simplest clustering algorithm
K-Means is the standard starting point for clustering. It is fast, interpretable, and works surprisingly well on a huge range of problems. The name tells you two things: K is the number of clusters you specify in advance, and "Means" refers to the centroid (average position) of each cluster.
Steps 2 and 3 repeat until the centroids stop moving. Convergence usually happens in 10 to 20 iterations. The final centroids define the cluster centres. Every new data point is assigned to whichever centroid it is closest to.
The algorithm is deceptively simple, but it makes one important assumption: clusters are roughly spherical and similarly sized. If your data has elongated, irregular, or very differently sized groups, K-Means will struggle. For those cases, alternatives like DBSCAN (which finds clusters of arbitrary shape) or Gaussian Mixture Models (which assign probabilities rather than hard labels) work better.
Choosing the right K
The biggest practical challenge with K-Means is that you have to specify K before you run the algorithm. How many clusters are there in your data? If you already knew that, you would not need clustering. You are trying to discover structure, not impose it.
The elbow method is the standard way to make this choice. You run K-Means with K ranging from 1 to 10 (or more), and for each K you record the inertia: the total squared distance from every point to its assigned centroid. As K increases, inertia always decreases. But the rate of decrease slows dramatically after the "true" number of clusters. You look for the elbow in the curve.
The curve drops steeply from K=1 to K=3, then starts to flatten. The elbow at K=3 suggests three is the natural number of clusters in this dataset. Adding more clusters beyond this point gives diminishing returns in how well the model fits the data.
The elbow is often a judgment call rather than a crisp mathematical answer. If the curve is smooth with no clear bend, the silhouette score is another metric worth checking. It measures how tightly each point fits within its own cluster compared to its distance from neighbouring clusters. Values close to 1 mean good clusters; values near 0 mean overlapping clusters.
Where clustering shows up in the real world
K-Means in Scikit-learn
The Scikit-learn API for clustering is the same pattern you have seen for everything else. Create the object, fit it, get the labels out. The main difference is that fit does not need a y (target) argument, because there are no labels.
import numpy as np import matplotlib.pyplot as plt from sklearn.cluster import KMeans from sklearn.preprocessing import StandardScaler # Simulate customer data: annual spend and visit frequency np.random.seed(42) spend = np.random.normal(loc=[500, 2000, 8000], scale=[100, 300, 800], size=(100, 3)).flatten() visits = np.random.normal(loc=[2, 8, 20], scale=[1, 2, 4], size=(100, 3)).flatten() X = np.column_stack([spend, visits]) # Scale the features (important for distance-based algorithms) scaler = StandardScaler() X_scaled = scaler.fit_transform(X) # Fit K-Means with 3 clusters kmeans = KMeans(n_clusters=3, random_state=42, n_init=10) labels = kmeans.fit_predict(X_scaled) # Summarise each cluster for k in range(3): mask = labels == k print(f"Cluster {k}: {mask.sum()} customers | " f"avg spend=${X[mask,0].mean():.0f} | " f"avg visits/yr={X[mask,1].mean():.1f}")
Cluster 1: 100 customers | avg spend=$7,984 | avg visits/yr=20.2
Cluster 2: 100 customers | avg spend=$1,997 | avg visits/yr=8.0
Three distinct customer segments emerge automatically: low-spend infrequent visitors, mid-range regular customers, and high-value loyal shoppers. No human had to categorise a single customer. The model found the natural groupings in the data.
K-Means uses Euclidean distance. If one feature is measured in thousands (annual spend) and another in single digits (monthly visits), the distance will be dominated by the large-scale feature. Always run StandardScaler first so each feature contributes equally to the distance calculation. Skipping this step is one of the most common mistakes beginners make with clustering.